Papers with large-scale multimodal pre-training
Mitigating Visual Knowledge Forgetting in MLLM Instruction-tuning via Modality-decoupled Gradient Descent (2025.findings-emnlp)
Copied to clipboard
Junda Wu, Yuxin Xiong, Xintong Li, Yu Xia, Ruoyu Wang, Yu Wang, Tong Yu, Sungchul Kim, Ryan A. Rossi, Lina Yao, Jingbo Shang, Julian McAuley
| Challenge: | Existing fine-tuning and continual learning methods compress visual representations and emphasize task alignment over visual retention. |
| Approach: | They propose a modality-decoupled gradient descent (MDGD) that regulates gradient updates to preserve effective rank of visual features and explicitly disentangles visual learning from task-specific alignment. |
| Outcome: | The proposed model reduces visual forgetting and improves visual retention . it disentangles visual learning from task-specific alignment and preserves effective rank . |
Modular and Parameter-Efficient Multimodal Fusion with Prompting (2022.findings-acl)
Copied to clipboard
| Challenge: | Recent research has made impressive progress in large-scale multimodal pre-training. |
| Approach: | They propose to use prompt vectors to align multimodal modalities by pretraining text inputs with prompts or embedding vectors. |
| Outcome: | The proposed method achieves comparable performance to several other multimodal fusion methods in low-resource settings. |